Genomics, Proteomics & Bioinformatics
Preprints posted in the last 90 days, ranked by how well they match Genomics, Proteomics & Bioinformatics's content profile, based on 188 papers previously published here. The average preprint has a 0.11% match score for this journal, so anything above that is already an above-average fit.
Luo, W.; Wu, R.; Peng, Z.; Tan, K.; Zhu, D.; Ouyang, X.; Xiao, Z. X.; Liu, Z.; Liu, H.; Chang, X.; Yin, Z.; Li, J.; Xinyu, Z.; Liu, X.; Liu, D.
Show abstract
The intermittent energy restriction (iER) represents an effective dietary strategy for improving metabolic diseases including metabolic dysfunction-associated steatotic liver disease (MASLD) and type 2 diabetes mellitus (T2DM), yet the underlying mechanisms remain elusive. In this study, we integrated human clinical data, mouse models, and in vitro experiments to investigate the role of iER in modulating the gut-liver axis in comorbid MASLD and T2DM. We demonstrate that an iER diet improves hyperglycemia, hepatic steatosis and decreases the abundance of gut pathogen Klebsiella pneumoniae, which is strongly associated with reductions in blood endotoxin, lipopolysaccharide (LPS) levels, suggesting a potential role of K. pneumoniae-derived LPS in mediating effects of the iER on hepatometabolic improvements. We confirm that K. pneumoniae-derived LPS exacerbates lipid accumulation and inflammation using an in vitro model. Mechanistically, we reveal a core target of protein lysine acetylation (Kac), hydroxyacyl-CoA dehydrogenase -subunit (HADHA) Lys353 in the liver of db/db mice through a multi-omics analysis. The iER decreases HADHA-K353 acetylation and enhances its enzyme activity. A Kac-mimicking mutation (K353R) increases its enzyme activity and stability, blocks its binding to the inflammasome adaptor ASC, and alleviates lipid accumulation and inflammation in K. pneumoniae-derived LPS induced in vitro model. This study provides novel insights into the potential benefits of the iER in comorbid MASLD and T2DM.
li, y.; Liu, Y.; wu, j.; liu, s.; lin, x.; guo, k.; yang, t.; feng, m.; zhang, h.; wang, x.; xing, w.; qian, s.; yang, r.; zhao, c.
Show abstract
BackgroundGray is one of the relatively rare coat colors in donkeys. The Hetian Gray donkey is a distinctive indigenous breed from the Xinjiang Uygur Autonomous Region of Northwestern China, characterized by progressive hair depigmentation with aging while retaining dark skin pigmentation. However, the genetic basis underlying this unique gray coat color phenotype remains unclear. ResultsTo elucidate the genetic basis, we conducted whole-genome resequencing of Gray and non-Gray donkeys. Genome-wide selection signature analyses identified a candidate region on chromosome 15. Subsequent fine-mapping using mass spectrometry-based genotyping of 42 loci refined the candidate interval and revealed a SNP within intron 2 of the ASIP gene, located in a genomic fragment with highly similar sequences, showing complete association with the gray coat color. Association analysis in an expanded population further confirmed a strong correlation between this variant and the gray phenotype. Gene expression analyses also supported the role of ASIP in regulating pigmentation in donkeys. ConclusionsThese findings identify a genetic determinant of gray coat color in donkeys and provide new insights into the molecular mechanisms underlying age-related depigmentation in domestic animals.
Sadhukhan, S.; Kumari, K.; Rout, P.; Panda, A. C.
Show abstract
HighlightsO_LIIdentified hundreds of potential chromatin-associated circRNAs in HEK293 cells, H9, and HeLa cells C_LIO_LIThe first report suggesting global interaction of circular RNAs with chromatin C_LIO_LIChromatin-associated circular RNAs interact with various RBPs involved in RNA splicing or processing C_LI Circular RNAs (circRNAs) have emerged as novel regulators of gene expression by interacting with various proteins and RNAs in a spatiotemporal manner. CircRNAs localized in the cytoplasm regulate mRNA translation or stability by binding to microRNAs and RNA-binding proteins (RBPs), while circRNAs in the nucleus regulate transcription and pre-mRNA splicing by associating with transcription factors and splicing factors. In this study, we sought to explore the interaction between circRNAs and chromatin. Analyzing published RNA-seq data from chromatin fractions identified hundreds of chromatin-associated circRNAs (cacRNAs) in various human cells. We validated the enrichment of a subset of circRNAs in the chromatin fraction and established the direct interaction of circDYNC1H1 and circKIF2C with chromatin in HEK293T cells. Furthermore, cacRNAs were found to interact with RBPs. Together, our research demonstrates the global association of hundreds of circRNAs with chromatin and expands our understanding of novel functional aspects of the circRNAs. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=186 SRC="FIGDIR/small/739476v1_ufig1.gif" ALT="Figure 1"> View larger version (69K): org.highwire.dtl.DTLVardef@1fc8831org.highwire.dtl.DTLVardef@516cdcorg.highwire.dtl.DTLVardef@1c20395org.highwire.dtl.DTLVardef@79313a_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOGraphical abstractC_FLOATNO C_FIG
Zhou, Y.; Huang, F.; Zhao, Y.
Show abstract
Tuberculosis remains a major global public health threat. While whole-genome sequencing has transformed our understanding of the causative agent, Mycobacterium tuberculosis (MTB), existing genomic databases are highly fragmented and often underrepresent structural variations (SVs). Furthermore, critical population-genetic statistics are rarely integrated with phylogenetic and geographic context, forcing researchers to reconcile separate datasets manually. To address this gap, we developed TBpop (https://tbpop.chinacdc.cn), an open-access, integrated population genomics portal. TBpop is built from 420 clinical MTB isolates selected from the first national drug resistance baseline survey in China. The portal integrates isolate metadata, pangenome categories, SNPs, SVs, IS6110 insertion sites, strain phylogeny, and gene-level statistics, and provides three interactive explorer modules: the Population Explorer, the Statistics Explorer, and the Variation Explorer. Additionally, a User Analysis module allows researchers to run population genetic workflows on their own alignments. TBpop provides an integrated platform for exploring genome plasticity, signatures of positive selection, and conservation patterns of functionally important genes in MTB.
Zhang, R.;Zhang, Y.;Zang, H.;Lou, J.;Li, Y.;Jiang, J.;Chen, D.;Yan, T.;Guo, R.
Show abstract
Melittin, the principal bioactive peptide of bee venom, exerts potent antitumor activity against hepatocellular carcinoma (HCC). However, the comprehensive transcriptomic alterations it elicits in hepatoma cells remain poorly characterized. Here, we present an integrated transcriptome dataset from melittin- and un-treated murine Hepa 1-6 hepatoma cells, encompassing messenger RNA (mRNA) and microRNA (miRNA) expression profiles. Cells were exposed to 4 g/mL melittin in serum-free DMEM for 20 min, and total RNA was subjected to ribosomal RNA-depleted strand-specific RNA sequencing on an Illumina NovaSeq6000 platform (paired-end 150 bp) and small RNA sequencing on an Illumina HiSeq2500 platform (single-end 50 bp). Raw data were processed using Cutadapt to remove adapters and low-quality reads, yielding clean datasets with Q20 [≥] 99.85%, Q30 [≥] 98.48%, and valid data ratios exceeding 85%. All raw and processed sequencing data are publicly available. This transcriptomic resource provides a valuable resource and basis for elucidating the regulatory networks underlying melittin-induced anti-hepatoma effects. DatasetThe dataset can be accessed through the National Genomics Data Center, China National Center website by searching with the BioProject accession number PRJCA065485 Reviewers may use this link for anonymous access during the review process. Direct URL to data: Genome Sequence Archive-CNCB-NGDC. Dataset LicenseCC BY 4.0
Lin, Y.; Chithravel, V.; Dai, J.; Liu, S.; Lubman, N. Y.; Lubman, D. M.
Show abstract
Hepatocellular Carcinoma (HCC) arising from Metabolic Dysfunction-Associated Steatotic Liver Disease (MASLD) is an increasing public health burden with high mortality, highlighting the need for improved early detection strategies. Current surveillance tools, including Alpha-fetoprotein (AFP) and ultrasound, lack sufficient sensitivity for early-stage HCC detection. We analyzed serum samples from 131 patients, including 58 with cirrhosis and 73 with MASLD-related HCC (42 early-stage, 31 late-stage), using an nLC-stepped HCD-PRM-MS/MS workflow for targeted N-glycome profiling of glycopeptides derived from haptoglobin and vitronectin. Combining targeted glycopeptides with AFP significantly improved HCC detection compared with AFP alone. The optimal panel for all HCC versus cirrhosis (AFP + VTNC_169_A2G2F0S1 + VTNC_242_A3G3F2S2) achieved an AUC of 0.859 and 76.7% sensitivity at 90% specificity. For early-stage HCC, AFP + HP_184_A3G3F1S3 + VTNC_169_A2G2F0S1 yielded an AUC of 0.890 with 66.7% sensitivity at 1% specificity. A SHAP-selected Gaussian Naive Bayes model based on seven molecular/glycopeptide features, without demographic variables, further improved performance, achieving ROC-AUC values of 0.9985 in training and 1.0000 in independent testing cohorts, with accuracies of 98.1% and 100.0%, respectively.
Wang, Q.; Wang, B.-Y.; Wilus, D.; Hua, X.
Show abstract
Periodontitis, a chronic inflammatory disease affecting approximately 40% of U.S. adults aged 30 years and older, is characterized by dysbiosis of the dental plaque microbiome. However, although scaling and root planing (SRP) is the cornerstone of periodontal treatment, its effects on the taxonomic composition and functional potential of the dental plaque microbiome remain incompletely understood. In this study, we used whole-metagenome shotgun sequencing to characterize taxonomic composition and functional potential in dental plaque microbiomes collected from 39 patients with Stage II or III generalized periodontitis before and 3-4 months after SRP. Consistent with clinical improvement, periodontal therapy significantly reduced bleeding on probing and plaque index. Whole-metagenome shotgun sequencing identified 3.18 million non-redundant genes and 12,353 microbial species across 78 samples, revealing increased gene and species richness after treatment, along with a significant restructuring of microbial community. Established periodontal pathogens, including Porphyromonas gingivalis and Tannerella forsythia, as well as the emerging pathogen Escherichia coli, decreased following treatment, whereas health-associated early colonizers, including multiple Actinomyces species and Streptococcus cristatus, increased. Functional annotation using the Carbohydrate-Active Enzymes (CAZy) database identified treatment-associated differences in several carbohydrate-active enzymes, including multiple glycosyltransferases, indicating remodeling of the predicted functional potential of the dental plaque microbiome. These findings demonstrate that successful SRP promotes coordinated taxonomic and predicted functional remodeling of the dental plaque microbiome and highlight the value of shotgun metagenomic sequencing for characterizing both taxonomic and functional recovery following periodontal therapy.
Yang, Y.; Zhang, N.; Li, T.; Wang, H.; Huang, X.; Ma, R.; Zhang, H.; Jing, X.; Di, R.; Xia, Q.; He, X.; Guo, X.; Zhang, X.; Jiang, Y.; Li, R.; Chu, M.; Liu, Q.
Show abstract
Seasonal breeding is a remarkable adaptive trait, but it constrains efficient production in the sheep industry. Recent studies have shown that seasonal breeding is associated with endogenous circannual rhythms, which are regulated in part by the circadian clock system. FBXL3, a pivotal component of the SCF (SKP1 - CUL1 - F-box) E3 ubiquitin ligase complex, is a known determinant of the mammalian circadian period. In this study, we identified a missense mutation, T183M, in FBXL3 through selective sweep analysis. The allele frequency of this mutation differed significantly between sheep breeds exhibiting year-round estrus and those showing seasonal breeding patterns. The association between the T183M mutation and seasonal breeding was further validated using an ovariectomized, estradiol-implanted sheep model. We then generated mice carrying the homologous T183M mutation and found that they exhibited significantly lengthened circadian periods, accompanied by reduced CRY1 expression and increased CLOCK expression. Co-immunoprecipitation assays confirmed that the mutation reduced the interaction between FBXL3 and CRY1. These findings demonstrate an evolutionarily conserved role of FBXL3 in the circadian clock system. We propose that the T183M mutation disrupts day-length recognition, thereby influencing seasonal estrus in sheep. Author summarySeasonal breeding limits sheep productivity and is regulated by circadian rhythms, yet the key genetic determinants remain poorly understood. Here, we identified an FBXL3 T183M missense mutation whose allele frequency differed markedly between year-round-estrous and seasonally breeding sheep, and validated its association with seasonal reproduction in a sheep population. Functional analyses in mutant mice showed that this variant lengthened the circadian period, disrupted the expression of core clock genes, and weakened the interaction between FBXL3 and CRY1. These findings suggest that the FBXL3 T183M variant impairs day-length perception, thereby modulating seasonal estrus in sheep. Our study reveals a conserved circadian mechanism underlying seasonal breeding and highlights FBXL3 T183M as a promising genetic target for improving reproductive performance in sheep.
Agrawal, A.; Kumar, S.; Vindal, V.
Show abstract
A protein whose removal or deletion causes significant disruption or collapse of a protein-protein interaction (PPI) network is referred to as a vulnerable protein. Such proteins may serve as valuable therapeutic or diagnostic targets in disease-associated networks. In this study, two PPI networks were constructed, one for HPV-positive and the other for HPV-negative head and neck squamous cell carcinoma (HNSCC), and the vulnerable proteins of these networks were identified by the node deletion approach. After analyzing the networks, 27 unique vulnerable proteins in HPV-positive and 72 unique vulnerable proteins in HPV-negative HNSCC were identified. Among them, one HPV-positive and seven HPV-negative HNSCC vulnerable proteins were further chosen by integrating multi-omics data. To exploit the vulnerabilities of these proteins, candidate synthetic lethal (SL) partners were predicted whose inhibition may selectively impair tumor survival. Subsequently, drug-gene interaction analysis was performed to identify inhibitors targeting the SL partners of these vulnerable proteins. Notably, in HPV-positive HNSCC, TOP2A, CHEK1, and CHEK2 genes were identified as SL partners of TTN, and their inhibitors were already clinically approved. While in HPV-negative HNSCC, ADA and MMP19 were identified as an SL partner of LMO7; TMEM45B, CDH3, and ELF3 genes were identified as an SL partner of CGN; and ZNF433 was identified as an SL partner of FLNC. However, MMP19, ZNF433, and TMEM45B inhibitors were not reported. Thus, these vulnerable proteins, including their SL partners, provide novel avenues to explore and develop more efficient and precise therapeutic and diagnostic strategies.
Rymbekova, A.; Kuhlwilm, M.
Show abstract
Archaic introgression has shaped the evolutionary history of Eurasian populations, yet Central Eurasian region remains understudied despite being at the crossroads of ancient human migration. Here, we analyzed the whole-genome data of five Central Eurasian (CE) individuals from Early Bronze Age (EBA) and five present-day CE individuals to characterize the archaic introgression landscape. We estimated that archaic introgression from Neanderthal and Denisovan archaic hominins comprises approximately 2.2% of the Central Eurasian genomes. Both amount and chromosomal distribution of archaic introgression remained largely unchanged between the EBA and present-day CE individuals. Putative introgressed fragments matching the Altai Neanderthal and the Altai Denisovan were retrieved. Our results suggest that while the archaic introgression levels seemingly remained stable over the past several thousand years, larger modern CE genomes panels will be required to fully characterize the genomic landscape of archaic ancestry in the region.
Wang, G.; Kubelt, C.; Smicius, R.; Zidane, K.; Rohrandt, C.; Brändl, B.; Wong, D.; Steiger, M.; Lum, A.; Evers, M.; Schmidt, N. O.; Pröscholdt, M.; Riemenschneider, M. J.; Kretzmer, H.; Synowitz, M.; Yip, S.; Vingron, M.; Müller, F.-J.
Show abstract
Copy number variations (CNVs) can serve as important clinical biomarkers for tumor classification and stratification. However, the utility of these CNV biomarkers for intraoperative tumor assessment within the timeframe of neurosurgical procedures has remained elusive due to the protracted duration of conventional CNV characterization methods. Here, we introduce CNVisor, a statistical framework for reliable and robust CNV detection from long-read sequencing, even under ultra-low coverage. Applied to neurosurgical tumor specimens, the proposed method enabled genome-wide CNV profiling and identified clinically relevant CNVs using roughly 60,000 reads within 20 minutes of sequencing. Integrating CNVisor with methylation-based classifiers can further reduce turnaround time and increase the accuracy of glioma subtype stratification. Together, these findings establish real-time CNV profiling using ultra-low coverage nanopore sequencing as a feasible strategy for intraoperative, genomics-informed assessment of CNS tumors.
Tang, R.; Liu, J.; Zhang, P.; Liang, X.
Show abstract
Background and objectiveGene regulatory networks are formed by complex regulatory relationships between transcription factors and their target genes. A systematic understanding of these regulatory relationships is crucial for deciphering the molecular mechanisms that underlie cell state transitions under physiological and pathological conditions. Single-cell expression data can reveal cell-type-specific transcriptional regulation, and computational methods have recently been developed to infer gene regulatory networks from single-cell transcriptomics and prior regulatory knowledge. However, existing methods could not explore the common and specific information in expression correlations and prior regulatory knowledge, which can adversely affect prediction performance. MethodsWe propose a novel method for inferring gene regulatory networks from single-cell RNA sequencing data. The proposed method consists of dual-channel graph neural networks and a weight-shared common graph neural network, enabling effective fusion of prior regulatory knowledge with gene co-expression patterns. Furthermore, we formulate a new computational framework built upon the proposed algorithm, which integrates differential gene expression profiles and regulatory changes to identify key regulators that distinguish different cell states. ResultsExperimental results demonstrate that our method significantly improves the accuracy of regulatory inference across multiple datasets, outperforming other state-of-the-art approaches. Our method also exhibits robustness to noise and missing data. Analysis of two single-cell expression datasets suggests that the proposed framework could help identify key regulators involved in tumor metastasis and drug resistance. ConclusionThese results indicate that the proposed method could advance the understanding of the biological mechanisms underlying diseases by reconstructing single-cell gene regulatory networks and identifying key regulators across different cell states.
Visser, C. d.; Rahm, L.; Lewerissa, E.; Mijdam, R.; Doornbos, C.; Huang, J.; O'Gorman, L.; Badmus, F.; van Karnebeek, C. D. M.; Faber, C. G.; Verhoeven, J.; van Bokhoven, H.; Kasri, N. N.; Lefeber, D.; 't Hoen, P. A. C.; van Gool, A. J.; Kulkarni, P.
Show abstract
Induced pluripotent stem cells (iPSCs) are widely used as patient-specific disease models, yet substantial unexplained variability in molecular and functional readouts limits their reliability. Here, we systematically investigated the sources of variation in iPSC-derived neurons for three rare genetic disorders: Myotonic Dystrophy Type 1, chromodomain-DNA-helicase-binding protein 2-related disorder and N-acetylneuraminic acid synthase deficiency. This was performed by profiling multi-omics layers: genomics, epigenomics, transcriptomics, proteomics, metabolomics and lipidomics. Our study found that clonal variability was comparable to inter-patient differences and that neuronal differentiation state and nutrient-driven metabolic activity emerged as dominant contributors to variability observed across omics layers. Clonal differences could partly be attributed to stochastic differences in DNA methylation established during reprogramming. By modeling and correcting the observed variation, we improved the detection of disease-associated molecular signatures. Our study provides guidelines for improved study design and data analysis to minimize variability, enabling robust biomarker discovery and reliable iPSC-based disease modeling.
Li, X.; Jiang, X.; Dong, Q.; Wu, J.; Li, Y.; Zhang, Y.; Zhong, L.
Show abstract
Background: Multiple myeloma (MM) progression is accompanied by remodeling of the bone marrow immune microenvironment. Local interactions among malignant plasma cells, stromal cells, myeloid cells, and immune cells not only support tumor cell survival, expansion, and immune escape, but are also closely associated with disease progression, therapeutic response, and clinical prognosis. Moreover, T cell exhaustion is a common T cells dysfunction in MM and limited efficacy of T cell-targeting therapies. However, the in situ organization and clinical significance of exhausted T cells in MM patients bone marrow remain insufficiently understood. Methods: In this study, we analyzed bone marrow Xenium 5K spatial transcriptomics data from control (Ctrl), monoclonal gammopathy of undetermined significance (MGUS), smoldering myeloma (SM), and MM samples. After canonical multi-sample integration and celltype annotation, we used Gaussian mixture model (GMM)-based spatial partitioning, and multilayer perceptron (MLP) machine learning for systematic characterization the T cell microenvironment in MM bone marrow. Results: Our results showed that exhaustion-like T cells increased during MM progression and formed spatially discrete T cell-enriched regions in the bone marrow, which we defined as exhaustion-like bone marrow T cell islands (eBM-TIs). These niches were mainly characterized by enhanced T cell-plasma cell communication associated with upregulated Galectin signaling. Pseudobulk analysis further showed enhanced IFN-related signaling in eBM-TIs, accompanied by upregulation of CXCR3 ligands such as CXCL9 and CXCL10, suggesting that the IFN-CXCL9/10 axis may contribute to T cell chemotaxis, maintenance of chronic inflammation, and formation of exhaustion-like states. By transferring spatial niche labels to scRNA-seq cohorts with available clinical staging information using MLP, we further found that the proportion of eBM-TI-like T cells was associated with higher disease risk and unfavorable prognostic outcomes. Conclusions: In summary, this study identifies eBM-TIs as a spatial niche in the MM bone marrow. These niches represent an important immune unit linking chronic inflammation, T cell exhaustion, and clinical risk, and may serve as a potential biomarker of MM disease progression.
Jiang, Y.; Luo, H.; Zheng, H.; Li, C.; Zan, X.; Xu, J.; Chen, Y.
Show abstract
Despite significant advancements in microsurgical techniques in recent years, the treatment and prognosis of craniopharyngiomas remain unsatisfactory. As a central nervous system tumor located adjacent to important brain structures such as the hypothalamus-pituitary axis and accompanied by a highly inflammatory microenvironment, the tumor heterogeneity and tumor microenvironment characteristics of papillary craniopharyngiomas (PCPs) remain unclear. In this study, we integrated multimodal single-cell and spatial profiling from PCP tissue and peripheral blood mononuclear cells (PBMCs) to elucidate the tumor heterogeneity and microenvironment characteristics of PCP. Our single-cell and spatial analyses defined four specific tumor cell states in PCP, representing specific transcriptional regulatory programs and spatial heterogeneity characteristics during tumor progression. By constructing a spatial niche composed of tumor, immune, and stromal cells, we analyzed the cellular and spatial ecosystem of PCP at multiple levels to further assess the communication relationships between different tumor cell states and microenvironment cells. This study established a multidimensional molecular atlas of PCP from the perspectives of cell state, spatial structure, and microenvironment interactions, providing a foundation for understanding its biological behavior and exploring new intervention strategies.
Goto, G.; Hanawa, D.; Naito, K.; Wang, Q. S.; Kanai, S.; Awaji, M.; Nishikawa, H.; Yui, H.; Nishitani, S.; Miyake, K.; Ooka, T.
Show abstract
Background: Large-scale biobanks have advanced genomic and epidemiologic research, but many rely on infrequent biological sampling and limited digital phenotyping. The Yamanashi Multi-omics Cohort (YMoC) was established to support longitudinal assessment of molecular, clinical, and behavioural changes in a screening-defined cohort of adults at elevated metabolic risk without diagnosed diabetes. Methods: YMoC is a longitudinal cohort of 215 adults aged 30-70 years in Yamanashi Prefecture, Japan, who met prespecified glycaemic eligibility criteria at health check-up, including fasting plasma glucose 100-125 mg/dL (5.6-6.9 mmol/L) and HbA1c <6.5%. Participants underwent three in-person visits over six months. Measurements include 75-g oral glucose tolerance testing with serial sampling, clinical biochemistry, anthropometry, liver elastography, and collection of blood, urine, stool, and saliva for multi-omics profiling. Between visits, participants wore a Fitbit Inspire 3 and completed daily app-based questionnaires using the Taohealth app. Current molecular data include genome-wide single nucleotide polymorphism array genotyping and longitudinal plasma proteomics in a subset. Conclusions: YMoC is designed to evaluate within-person molecular and phenotypic trajectories in a screening-defined metabolic-risk cohort. The cohort provides a dense longitudinal resource linking clinical assessments, biospecimens, omics assays, and digital phenotyping, including analyses of insulin-resistance-related markers such as homeostasis model assessment of insulin resistance (HOMA-IR).
Peng, C.; Schreiber, H.; Zhang, C.; Liu, Q.; Hultgren, S. J.; Freddolino, L.
Show abstract
The rapid advancement of high-throughput sequencing technologies has vastly increased the number of known protein sequences, but the experimental characterization of their structures and functions lags behind. This gap in knowledge impedes our understanding of biological mechanisms of these proteins, hinders the interpretation of high-throughput experiments, and exposes a significant challenge in modern biology: deducing the structural and functional information of proteins based on their sequences. Most computational approaches rely on homology with well-annotated proteins, yet many proteins lack identifiable homologues, reducing the power of this approach. Here, we integrated cutting-edge protein structure and function prediction methods to develop a complete sequence-structure-function pipeline that predicts structures and functions based on primary sequences. We applied this pipeline to predict the structure and function of all proteins in Escherichia coli UTI89, a model strain of uropathogenic E. coli. Based on the predicted functions, we performed enrichment analysis on the whole genome and revealed the possible roles and related biological mechanisms of poorly annotated proteins in this organism. Moreover, the performance of our pipeline was further validated through detailed case studies of the UTI89_C0931 and ybtS genes. Finally, we compiled the UTI89 structure and function database (https://seq2fun.dcmb.med.umich.edu/UTI89), offering it as a community resource to aid researchers in elucidating the roles of unannotated proteins in uropathogenic E. coli. This database aims to bridge critical knowledge gaps in microbial pathogenicity and resistance, enhancing our capacity to tackle emerging health threats.
Zhou, Y.;Jin, S.;Zhong, J.;Xiao, X.;Ding, M.;Zhao, L.;Guo, Z.
Show abstract
Tomato yellow leaf curl virus (TYLCV) is a devastating viral pathogen threatening agricultural crops globally. In this study, we identified a novel TYLCV isolate (TYLCV-YN6244), which caused viral epidemic in resistant tomato cultivars at Yuanmo county, Yunnan Province of China. We determined the complete genome of TYLCV-YN6244 and found it encoded six viral proteins characteristic of Geminivirus. We identified its V2 protein as a potent viral suppressor of RNA silencing (VSR), and generated infectious clone of wildtype TYLCV-YN6244, or V2-defective TYLCV-YN6244 (TYLCV-YN6244-{Delta}V2) in which V2 was deleted. Both of infectious clones were capable of systemically infecting tobacco and tomato. However, TYLCV-YN6244 but not TYLCV-YN6244-{Delta}V2 could cause disease symptoms in wildtype tobacco or tomato plants, and viral accumulation was drastically reduced in plants infected with TYLCV-YN6244-{Delta}V2 compared to TYLCV-YN6244 while the efficiency of virus-derived small interfering RNAs (vsiRNAs) biogenesis was conversely increased in plants infected with TYLCV-YN6244-{Delta}V2. Surprisingly, small RNA profiling indicated that 21nt and 22nt rather than 24nt vsiRNAs were predominantly produced in tomato plants infected with either TYLCV-YN6244 or TYLCV-YN6244-{Delta}V2. Furthermore, transcriptome analyses revealed that TYLCV-YN6244 or TYLCV-YN6244-{Delta}V2 infection differentially modulated metabolism and defense-related pathways in tomato, probably underlying distinct viral pathogenicity and disease symptoms induced in plants. Overall, our research not only identified a novel pathogenic TYLCV isolate but also characterized molecular biology and host response in tomato with infectious clones firstly developed, with implications in untangling virus-host interaction for developing novel resistance in crop tomato.
Khandelwal, S.; Jarvis, N.; Zhan, J.
Show abstract
Glioblastoma (GBM) is a highly aggressive brain tumor with an extremely poor 5-year survival rate of 6.9%, largely attributable to the lack of reliable biomarkers. While competing endogenous RNA (ceRNA) and copy number variation (CNV) analyses offer unique biomarker identification potential, current approaches neglect the integration of multiple regulatory mechanisms for biomarker detection. To address this limitation, we applied relational graph convolutional networks (RGCNs) to ceRNA and CNV knowledge graphs through a novel late fusion ensemble architecture. The proposed architecture outperformed baseline models and identified five novel biomarkers, including hsa-miR-196a and hsa-miR-224. Kaplan-Meier survival analysis and Cox regression indicated that the identified genes hold significant prognostic and diagnostic power. The early stratification of the Kaplan-Meier curves indicates the potential these genes hold for patient survival prediction. The results illustrate that a late fusion RGCN ensemble effectively captures complex gene interactions, overcoming limitations of existing models and providing a framework for biomarker discovery. The novel biomarkers serve as prospective targets for future GBM therapeutic development and candidates for non-invasive diagnostic assays.
Zhang, Y.-F.; Xu, Z.-h.; Gao, C.-x.; Duan, S.-Y.; Li, G.; Xu, C.; Lu, H.-M.
Show abstract
The attention mechanism offers the possibility for data-driven discovery of biological principles. However, for important protein families such as human olfactory receptors, the extent to which attention can associate with biologically meaningful key regions lacks systematic validation. In this study, using human olfactory receptors (ORs) as a model, we constructed CrossVOI, a VOC-OR interaction prediction framework based on protein language models and cross-attention, achieving predictive performance superior to existing methods. Furthermore, we systematically analyzed the attention distributions of CrossVOI and found that attention not only focused on ligand-binding interfaces and evolutionarily conserved sites, but also to some extent identified certain dynamically regulated regions. In summary, we propose CrossVOI, currently the best-performing framework for VOC-OR interaction prediction, and analyze the interpretability of the attention mechanism for human ORs. This study provides insights into the interpretability of protein function prediction methods and is expected to contribute to the exploration of attention mechanisms in biological mechanisms, and provide assistance for large-scale screening and mechanistic analysis of olfactory receptors.